Papers with ROUGE metric

12 papers
Global Voices: Crossing Borders in Automatic News Summarization (D19-54)

Copied to clipboard

Challenge: a crowd-sourced dataset is needed to evaluate cross-lingual summarization methods . human-written summarizing is expensive and difficult to design for humans .
Approach: They construct a multilingual dataset for evaluating cross-lingual summarization methods . they use social-network descriptions of news articles to extract evaluation data .
Outcome: The proposed dataset compares a translate-then-summarize approach with baselines in 15 languages.
SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at Scale (2022.emnlp-main)

Copied to clipboard

Challenge: Recent work has introduced an important yet relatively under-explored NLP task called Semantic Overlap Summarization (SOS) that entails generating a summary from multiple alternative narratives which conveys the common information provided by those narratives.
Approach: They propose to use a sentence-level precision-recall style automated evaluation metric to evaluate a new NLP task called Semantic Overlap Summarization (SOS) they propose to employ the popular ROUGE metric and use it to compare the two tasks.
Outcome: The proposed metric yields higher correlation with human judgment and higher inter-rater agreement compared to the existing metric.
OTExtSum: Extractive Text Summarisation with Optimal Transport (2022.findings-naacl)

Copied to clipboard

Challenge: Extractive text summarisation aims to select salient sentences from a document to form a short yet informative summary.
Approach: They propose to formulate extractive text summarisation as an Optimal Transport (OT) problem and use it to obtain an optimal summary that minimises the transportation cost to a given document.
Outcome: The proposed method outperforms state-of-the-art methods and learning-based methods on multiNews, PubMed, BillSum, and CNN/DM datasets.
Evaluating Multiple System Summary Lengths: A Case Study (D18-1)

Copied to clipboard

Challenge: Practical summarization systems are expected to produce summaries of varying lengths, per user needs.
Approach: They propose to use ROUGE metric to evaluate system summaries of multiple lengths.
Outcome: The evaluation protocol in question is competitive, the authors show . they found that the evaluation protocol is competitive with existing benchmarks.
Multi-Reward Reinforced Summarization with Saliency and Entailment (N18-2)

Copied to clipboard

Challenge: Abstractive text summarization is the task of compressing and rewriting a long document into a short summary while maintaining saliency, directed logical entailment, and non-redundancy.
Approach: They propose a novel reward function for ROUGESal and Entail to improve abstractive summarization . they use a coverage-based reward function to combine ROUGE and En Tail .
Outcome: The proposed method achieves state-of-the-art results on CNN/Daily Mail dataset and strong improvements in a test-only transfer setup on DUC-2002.
Revisiting Automatic Evaluation of Extractive Summarization Task: Can We Do Better than ROUGE? (2022.findings-acl)

Copied to clipboard

Challenge: Existing methods to evaluate text summarization tasks using ROUGE have been criticized for lack of semantic understanding.
Approach: They propose a semantic-aware metric for extractive summarization task that is semantic-based . they use CNN/DailyMail dataset to study the new metric .
Outcome: The proposed metric is semantic-aware and shows higher correlation with human judgement and yields a large number of disagreements with the original ROUGE metric.
An Anchor-Based Automatic Evaluation Metric for Document Summarization (2020.coling-main)

Copied to clipboard

Challenge: Existing reference-based evaluation metrics such as ROUGE have their own drawbacks.
Approach: They propose a protocol for a reference-based automatic evaluation metric that requires the endorsement of source document.
Outcome: The proposed metric is anchored on source document and has higher correlation with human judgments.
Cross-lingual Evaluation of Multilingual Text Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for multilingual text generation are limited by language and data leakage.
Approach: They propose an annotation-free cross-lingual evaluation protocol for multilingual text generation . they first generate English references from the translated non-English inputs into English .
Outcome: The proposed protocol shows a high correlation to the reference-based ROUGE metric in four languages on news text summarization.
Semantic Overlap Summarization among Multiple Alternative Narratives: An Exploratory Study (2022.coling-1)

Copied to clipboard

Challenge: Existing tasks for summarizing multiple alternate narratives with different perspectives are under-explored.
Approach: They propose a task which entails generating a single summary from multiple alternative narratives . they use a web-based dataset and human annotations to evaluate the task .
Outcome: The proposed task is based on a novel dataset and human annotations.
SumeCzech: Large Czech News-Based Summarization Dataset (L18-1)

Copied to clipboard

Challenge: Summarization of documents is a well-studied NLP task, but only a few datasets are available for Czech.
Approach: They propose to use a Czech news-based summarization dataset to evaluate document summarizing . they propose a language-agnostic variant of the ROUGE metric to enable automatic evaluation .
Outcome: The proposed dataset contains more than a million Czech news articles . the proposed approach is strong abstractive and language-agnostic .
A Summarization Dataset of Slovak News Articles (2020.lrec-1)

Copied to clipboard

Challenge: a number of studies on document summarization have focused on the English language . however, most of the work on this task is done on English datasets .
Approach: They propose to use a news site's ROUGE metric to adapt it to Slovak texts . they propose to introduce a large-scale news-based summarization dataset .
Outcome: The proposed approach is better suited for Slovak texts than the dominant ROUGE metric.
NumHG: A Dataset for Number-Focused Headline Generation (2024.lrec-main)

Copied to clipboard

Challenge: a lack of fine-grained annotations for accurate numeral generation in headlines is a major roadblock . a new dataset, the NumHG, provides over 27,000 annotated numeral-rich news articles for detailed investigation .
Approach: They propose a dataset that provides annotated numerals for headline generation . they evaluate five well-performing headline-generation models using human evaluation .
Outcome: The proposed dataset provides annotated numeral-rich news articles for detailed investigation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations